Chinese HowNet-Based Multi-factor Word Similarity Algorithm Integrated of Result Modification

نویسندگان

  • Benbin Wu
  • Jing Yang
  • Liang He
چکیده

In this paper, we firstly describe a novel approach to calculate the Chinese sememe similarity based on the HowNet hierarchical sememe tree. When we calculate the sememe similarity, we not only take Semantic Distance, Node Depth and Semantic Coincidence Degree into consideration, but also propose two impact factors named Node Environment Dense (NED) and Node Layer Ratio (NLR) to optimize the calculation process. Secondly, quite a few words described by identical concept definition in HowNet should have a certain discrimination according to human perception, so we propose a hybrid modification algorithm integrated of TongYiCi CiLin (hereinafter, CiLin) to deal with this case. Experiment results of the HowNet-based multi-factor similarity hybrid algorithm shows that this approach improves the similarity of independent sememe words and the words having identical concept descriptions in HowNet, while no large bias influence on the similarity of other words.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

A Word Similarity Algorithm with Sememe Probability Density Ratio Based on HowNet

The study on word similarity computation plays an important role in natural language processing (NLP). Recently the algorithm based on HowNet is widely used and proves to work well in Chinese word similarity computation. However, the relationship between the number of brother nodes and the fineness of the hierarchy is not considered. This paper investigates the ratio of two words on the brother...

متن کامل

An Unsupervised Approach to Chinese Word Sense Disambiguation Based on Hownet

The research on word sense disambiguation (WSD) has great theoretical and practical significance in many fields of natural language processing (NLP). This paper presents an unsupervised approach to Chinese word sense disambiguation based on Hownet (an electronic Chinese lexical resource). In our approach, contexts that include ambiguous words are converted into vectors by means of a second-orde...

متن کامل

The Research of Chinese Words Semantic Similarity Calculation with Multi-Information

Text similarity has a relatively wide range of applications in many fields, such as intelligent information retrieval, question answering system, text rechecking, machine translation, and so on. The text similarity computing based on the meaning has been used more widely in the similarity computing of the words and phrase. Using the knowledge structure of the and its method of knowledg...

متن کامل

Semantic Similarity Calculation of Chinese Word

This paper puts forward a two layers computing method to calculate semantic similarity of Chinese word. Firstly, using Latent Dirichlet Allocation (LDA) subject model to generate subject spatial domain. Then mapping word into topic space and forming topic distribution which is used to calculate semantic similarity of word(the first layer computing). Finally, using semantic dictionary"HowNet" to...

متن کامل

Word Similarity Computing Based on Hybrid Hierarchical Structure by HowNet

Word similarity computing is one of the most important and fundamental task in the field of natural language processing. Most of word similarity methods perform well in synonyms, but not well between words whose similarity is vague. It confronts the challenge of how to overcome this problem. An approach is proposed to compute Chinese word similarity based on hybrid hierarchical structure by How...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2012